Papers by Flor Miriam Plaza-del-Arco

11 papers
Countering Hateful and Offensive Speech Online - Open Challenges (2024.emnlp-tutorials)

Copied to clipboard

Challenge: a comprehensive understanding of the field is needed to maintain respectful and inclusive online environments.
Approach: This tutorial aims to provide attendees with a comprehensive understanding of the field by delving into essential dimensions such as multilingualism, counter-narrative generation, a hands-on session with one of the most popular APIs for detecting hate speech, fairness, and ethics in AI, and the use of recent advanced approaches.
Outcome: This tutorial aims to provide attendees with a comprehensive understanding of the field by delving into essential dimensions such as multilingualism, counter-narrative generation, a hands-on session with one of the most popular APIs for detecting hate speech, fairness, and ethics in AI, and the use of recent advanced approaches.
Responsible Evaluation of AI for Mental Health (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches to evaluating AI tools in this domain remain fragmented and inconsistent.
Approach: They propose a taxonomy of AI mental health support types that integrates clinical soundness, social context, and equity to provide a structured basis for evaluation.
Outcome: The proposed framework integrates clinical soundness, social context, and equity, providing a structured basis for evaluation.
Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions (2024.lrec-main)

Copied to clipboard

Challenge: Emotion analysis (EA) is a rapidly growing field in natural language processing . there is no consensus on scope, direction, or methods for EA .
Approach: They review 154 relevant NLP papers on emotion analysis from the last decade . they ask: how are EA tasks defined in NLP? what are the most prominent emotion frameworks and which emotions are modeled?
Outcome: The authors examine 154 relevant NLP papers on emotion analysis from the last decade . they find that there is no consensus on scope, direction, or methods .
MentalRiskES: A New Corpus for Early Detection of Mental Disorders in Spanish (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on the prevalence of mental disorders on the Web are limited to the English language.
Approach: They propose to use user messages posted on Telegram groups to annotate the corpus for natural language processing and to conduct experiments on text classification and regression.
Outcome: The proposed corpus contains over 1,300 subjects with more than 45,000 messages posted in different public Telegram groups.
Seeing Race, Feeling Bias: Emotion Stereotyping in Multimodal Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Emotion stereotypes are also tightly tied to race and skin tone, but previous studies have overlooked this dimension.
Approach: They propose a multimodal study of racial, gender, and skin-tone bias in emotion attribution . they evaluate four open-source MLLMs using 2.1K emotion-related events .
Outcome: The proposed study examines four open-source MLLMs using 2.1K emotion-related events paired with 400 neutral face images across three different prompt strategies.
SHARE: A Lexicon of Harmful Expressions by Spanish Speakers (2022.lrec-1)

Copied to clipboard

Challenge: Using natural language processing, offensive comments can be created by composition of words.
Approach: They propose to use a lexical resource with 10,125 offensive terms and expressions collected from Spanish speakers to retrieve the vocabulary.
Outcome: The proposed resource has 10,125 offensive terms and expressions and is used to identify spans in Spanish.
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation benchmarks for large language models lack annotations that justify moral classifications and focus on English constrain moral reasoning across diverse cultural settings.
Approach: They propose a multilingual benchmark dataset for evaluating moral reasoning of large language models . it includes 3,000 tweets annotated with binary hate speech labels, moral categories and rationales .
Outcome: The proposed dataset shows a misalignment between LLM outputs and human annotations in moral reasoning tasks.
La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America (2025.acl-long)

Copied to clipboard

Challenge: La Leaderboard is the first open-source leaderboard to evaluate generative Large Language Models (LLMs) in languages and language varieties of Spain and Latin America.
Approach: They propose to use La Leaderboard to evaluate generative Large Language Models in Spanish and Latin America.
Outcome: La Leaderboard is the first open-source leaderboard to evaluate generative LLMs in languages and language varieties of Spain and Latin America.
Your Mileage May Vary: How Empathy and Demographics Shape Human Preferences in LLM Responses (2025.findings-emnlp)

Copied to clipboard

Challenge: large language models (LLMs) increasingly assist subjective decision-making . prior work uses aggregate human judgments, but demographic variation and its linguistic drivers remain underexplored.
Approach: They analyze how demographic background and empathy level correlate with LLM-generated dilemma responses . they also identify markers that predict group-level differences .
Outcome: The authors show that demographic background and empathy level correlate with LLM preferences . their findings highlight the need for demographically informed LLM evaluations.
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluations of Large Language Models (LLMs) rely on a single large model to score outputs from other LLMs, but this is prone to intra-model bias and many tasks may be too subjective for a one model to judge fairly.
Approach: They propose a language model council where a group of LLMs collaborate to create tests, respond to them, and evaluate each other’s responses to produce a ranking in a democratic fashion.
Outcome: The proposed model produces rankings that are more separable and robust than any individual LLM judge.
Natural Language Inference Prompts for Zero-shot Emotion Classification in Text across Corpora (2022.coling-1)

Copied to clipboard

Challenge: Existing models for textual emotion classification depend on domain and application scenario and need to be predefined . a natural language inference model with a flexible set of labels is difficult to develop .
Approach: They propose to use the paradigm of zero-shot learning as a natural language inference task to generate a model with a flexible set of labels.
Outcome: The proposed model is more robust across corpora than individual prompts and shows similar performance to the best prompt for a particular corpus.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations